Papers by N J Karthika
Multilingual Tokenization through the Lens of Indian Languages: Challenges and Insights (2026.findings-acl)
Copied to clipboard
Maharaj Brahma, N J Karthika, Rajat Verma, Nagasai Saketh Naidu, Rohit Saluja, Maunendra Sankar Desarkar, Ganesh Ramakrishnan
| Challenge: | Existing tokenizers are often skewed towards high-resource languages limiting their effectiveness for linguistically diverse and morphologically rich languages. |
| Approach: | They evaluate multilingual tokenization across 17 Indic languages spanning 11 scripts and two language families. |
| Outcome: | The proposed method improves tokenization quality and vocabulary size in 17 languages . poor tokenization can lead to increase in sequence lengths, fragment meaningful units, weaken model's ability to capture linguistic structure and semantics. |
LexGen: Domain-aware Multilingual Lexicon Generation (2025.acl-long)
Copied to clipboard
Ayush Maheshwari, Atul Kumar Singh, N J Karthika, Krishnakant Bhatt, Preethi Jyothi, Ganesh Ramakrishnan
| Challenge: | Lexicon generation is a key task in specialized domains due to infrequent usage of terms . a new model is proposed to generate dictionary words for 6 Indian languages . |
| Approach: | They propose a model to generate dictionary words for 6 Indian languages in the multi-domain setting. |
| Outcome: | The proposed model generalizes to unseen domains and unsealed languages. |